Skip to content

feat(collaboration): wake the requesting lead once a delegated result is accepted - #5304

Open
songoow wants to merge 25 commits into
loopx-project:mainfrom
songoow:codex/delegation-wake-lead
Open

songoow wants to merge 25 commits into
loopx-project:mainfrom
songoow:codex/delegation-wake-lead

Conversation

@songoow

@songoow songoow commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Why

Roadmap R3: automatic return should continue the lead without owner polling. When a delegated member's result became accepted, nothing scheduled the requester's next bounded Turn. The App's return service only appends a follow-up message and never starts a Turn, and delegation routes are peer routes that do not even receive that message. A Goal Chat LoopX lead therefore saw a result only if it happened to call read/wait during its own Turn.

Seam decision

The wake is emitted on the first transition to accepted, not on adoption. Adoption runs inside the lead's own Turn, where planChatMode refuses a new Turn because one is already active.

What changed

  • Intent. transitionDelegationObservation returns a wake_intent (requester identity, operation, request and accepted artifact digests hashed into intent_id) only on a non-accepted to accepted transition; accepted to accepted emits nothing.
  • Persistence. _observe stores the intent as wake: {..., state: "pending"} in the same write that records accepted, so result and intent are crash-atomic. An in-Turn read/wait/adopt that observes the accepted result marks it observed_in_turn.
  • Admission. LoopXMode.wake and a wake operation in planChatMode reuse the resume rules plus mode enabled and not paused, and accept only a host origin. On admission it submits /goal resume with client_turn_id = "wake-" + intent_id[:32], so turn_for_client makes a retry idempotent. The session lock is taken before the record lock, the same order the in-Turn tool uses.
  • Pump. A delegation wake pump runs beside the return service in the Chat server. It resolves the owner from the intent's requester, skips an intent stored under another requester's address, and isolates per-record failures.
  • Receipts. The result and the wake are separate facts: woken (session, Turn, created), pending (lead_turn_active, lead_paused, allowance_exhausted) or refused (goal_stopped, lead_unbound, binding_revoked, native_goal_complete, native_goal_absent, wake_identity_conflict, no_wake_owner). A wake never unpauses the lead, never starts a native Goal and never raises the allowance.
  • Placement. The pump lives in chat_loopx_mode.py beside the owner it drives. An earlier revision added a new top-level module, which broke the top-level module budget.
  • Census and docs. The registry I/O manifest is regenerated for the pump's Goal context read. goal-chat-continuation.md and local-delegation.md describe the receipts and limits in English and Chinese.

Behavior change disclosure

A running Chat service now starts one Turn for an enabled, unpaused Goal Chat LoopX lead after its delegated result is accepted. Nothing changes for leads outside Chat LoopX mode (no_wake_owner, the scheduler deadline recheck remains their continuation), for paused leads, or when the Chat service is not running.

Checks run on this head

Check Result
pytest tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py tests/test_chat_delegation_journey.py tests/test_collaboration_mcp.py tests/test_local_delegation.py tests/test_delegation_inventory.py 55 passed
New wake tests alone 11 passed; mutating both accepted-status guards fails 3, removing the existing-Turn recovery fails 2
node --test tests/control_plane_ts/delegation.test.ts tests/control_plane_ts/chat_mode.test.ts 21 passed, 0 failed
npm run typecheck:control-plane ok
tests/architecture/test_top_level_module_budget.py, import boundaries, registry census passed
examples/docs-governance-smoke.py, examples/semantic-vocabulary-drift-smoke.py ok
loopx canary premerge --from-git-diff passed: tier=standard, changed_files=13, surfaces=control_plane/docs_project_content/public_boundary/python; selected=13, failures=0

Not run: a live model Turn started by a wake, the packaged App and Lark. The wake starts a Turn through the existing submit_turn; the App shows it as an ordinary LoopX execution Turn.

Coordination with other open PRs

Rollback

Revert the PR, or stop the pump alone: pending intents stay pending and wake nobody. Receipts already written remain readable facts on the operation record.

Bounded future-facing refactor

Applied: the pump sits inside the owner it drives instead of a new top-level module. Deferred: converging the return service and the wake pump into one Chat-host tick, which would change the return service's cadence and belongs with R3's return-lifecycle unification.

This is a control-plane change; it is left for maintainer review and merge.

🤖 Generated with Claude Code

songoow and others added 12 commits September 29, 2026 12:31
The transition to accepted is the one durable moment a requester can be
continued without polling. transitionDelegationObservation now returns a
wake_intent beside the accepted status: schema, an intent id derived from
the requester, operation, request and accepted artifact digests, and the
requester identity. accepted -> accepted stays an idempotent readback and
emits nothing; rejection wakes nobody. The intent grants no Turn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…in-Turn

The accepted observation now stores row["wake"] = {intent, state: pending}
in the same atomic write that records accepted, so a result and its wake
receipt cannot diverge. record_wake settles a pending intent under the
same lock adopt_result uses and never touches a non-accepted row; the
Chat LoopX tool marks the intent observed_in_turn whenever the lead sees
an accepted result inside its own Turn, so no redundant wake follows.
read_delegation exposes the receipt as a fact distinct from the result.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…owner

planChatMode gains a host-only "wake" operation that reuses the owner
resume facts (Goal active, no running Turn, registered coordinator, valid
bindings, resumable native Goal, allowance above consumed tokens) plus
mode enabled and not paused, and returns a typed outcome instead of an
error: admitted, pending (lead_turn_active, lead_paused,
allowance_exhausted) or refused (goal_stopped, no_wake_owner,
lead_unbound, binding_revoked, native_goal_complete, native_goal_absent).

ChatLoopXMode.wake takes the session lock before the record lock, asks the
rule, and submits one idempotent "/goal resume" Turn keyed by the intent
id, so a crash between submit and receipt recovers through turn_for_client
with created=false. apply rejects the host-only operation. The wake Turn
tells the owner why it started and tells the lead which operation to
read; it never unpauses the lead or starts a native Goal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
DelegationWakeService scans conversations that configured a coordinator,
maps their observed operations (loopx_deliveries) to the requester-scoped
records, and asks ChatLoopXMode.wake to settle each pending intent under
the same record lock adopt_result uses. It prefers a conversation that is
still enabled over one that exited, changes a pending receipt only when
its reason changes, and skips records held by a worker until the next
tick. Requesters outside a Chat LoopX conversation are never visited.
Closing the Chat server stops the pump; intents left pending are the
rollback state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…ting

The wake rule checked the native status before the running-Turn fence, so
an owner start that was still activating (Turn active, native "absent" or
the previous run's "complete") refused the wake terminally. The fence now
comes first; native complete/absent are refused only when no Turn runs.
The rule test covers every admitted, pending and refused outcome, and that
the host origin admits only wake.

The intent uses schema_version like every other collaboration contract and
requires the requester goal reference to be an object or null.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The pump visited only operations listed in a conversation's
loopx_deliveries, so requesters outside Chat LoopX mode stayed pending
forever and a lead with more than twenty operations lost wakes to the
bounded list. It now scans this runtime's operation records and resolves
the owner from the intent's requester identity, preferring a conversation
that can still continue and then the one that observed the operation; a
requester with no such conversation is refused with no_wake_owner. An
intent must name the requester that owns its storage address, and one
failing record no longer aborts the tick.

Receipts share one constructor that keeps only the typed intent and the
current state's facts. An existing wake-* Turn is recorded only when it
carries this intent; otherwise the wake is refused as
wake_identity_conflict rather than claiming another Turn. The Goal
context is read only after admission, so a removed Goal refuses instead
of raising. Adopting a result also retires the consumer's wake.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The typed rule accepted origin "web" for a wake, so the host-only
guarantee rested on the Python apply guard alone. Each origin now maps to
exactly one operation family: the host may only wake, and the owner's web
origin may never request one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The pump was a new top-level module, which raised loopx/ from 147 to 148
top-level modules and failed the module budget. It only drives
LoopXMode.wake, so it now lives beside that owner in chat_loopx_mode. The
chat_runtime and collaboration_mcp dependencies are imported inside the pump
because chat_runtime imports this module. No behavior change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Expectations follow the wake contract: an accepted result with a pending
intent continues its requester's lead exactly once; a second tick, a crash
after submit, a client id owned by another Turn, an active or paused lead, a
stopped Goal, a requester without a lead conversation, an intent stored under
another requester and a non-accepted result each leave the matching receipt
without starting a Turn. Mutating both accepted-status guards or the
existing-Turn recovery fails these tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Goal Chat continuation now lists the wake receipt states and reasons and the
limits (no unpause, no new native Goal, no higher allowance, requires the
Chat service). The shell entrypoint note points to it. English and Chinese
sections stay aligned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The Chat server reads the registry to build the Goal context for a
host-initiated wake Turn, so that read is a new codec_api site and the
server's other sites moved. Regenerated with
scripts/generate_project_registry_io_manifest.py; no unclassified sites.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Regenerated with scripts/generate_project_registry_io_manifest.py after
rebasing onto main. Site ids and classifications are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow
songoow force-pushed the codex/delegation-wake-lead branch from 4a72d07 to 3f99898 Compare September 29, 2026 16:32

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English verdict: REQUEST_CHANGES — persisted queued Turns are mistaken for completed wake dispatch, and an unpinned conversation selector permits duplicate dispatch after receipt loss.

Reviewed exact head: 3f99898d65d2d0c880f56797d1e46f35efab87e2, against immutable merge base 9c29941559cff92f675c37cc80c4cf87a23f8448. This review follows the current LoopX PR-review capability execution contract.

[P1] queued Turn 存在不等于唤醒已派发;启动失败后会永久跳过恢复

位置:chat_loopx_mode.py:494–504。native admission 先持久化 queued Turn,再启动 adapter。若 adapter 在这两个阶段之间失败,intent 仍 pending,但下次 tick 只检查 Turn 的身份是否匹配,便直接写 woken,没有重新进入原生 dispatch/recovery。

独立复现使用真实 ChatRuntimeController.submit_turn、TS acceptance 和 durable Chat store,只在 adapter/worker 传输边界注入故障:第一次 adapter 初始化失败,Turn 为 queued、started_at=None、dispatch=0;第二次 tick 后 adapter 尝试次数仍为 1、dispatch 仍为 0,wake 却是 woken。后续 pump 不再处理它,结果无法触发承诺的后续 Turn。现有测试对 submit_turn 的替身不能证明这段恢复。

最小修复是复用既有原生 acceptance/dispatch 恢复 owner,而非把“Turn 存在”当作派发完成。chat_turn_acceptance.ts:453 的 dispatchFor 已区分 queued、running 等状态。请覆盖 acceptance 后 adapter 失败、准备 capsule 后失败、queued/running/terminal 重试和 receipt 丢失,并验证只恢复原 Turn、不另建 Turn,同时保留 pause/revocation 边界。

[P1] 每次重选会话会破坏原会话返回和跨重试的单次派发

位置:chat_loopx_mode.py:923–934。selector 按 enabled/usable 优先、其次 operation 观测、最后更新时间选择同一 Goal/Agent 的会话,但 intent 没有固定原始 conversation;而 turn_for_client 的去重范围是 session,不是整个 intent。

独立复现一:原会话拥有该 operation,随后退出模式;新建同一 Goal/Agent 的 enabled 会话。pump 向新会话提交,而不是保留原会话的关闭边界。复现二通过真实 native acceptance:首 tick 已在原会话接受并派发 Turn,仅 wake receipt 写入注入一次 OSError;随后原会话 disabled,创建新的 enabled 会话;重试产生第二个 session 的新 Turn 和第二次 dispatch,最终回执指向替代会话。两个调用使用完全相同 intent/client id。这里只替换 worker 传输以避免运行模型;观测到的是两次真实 admission 和两次派发请求,不声称已花费模型额度。

请复用既有原会话返回关系,在 intent/恢复记录中稳定绑定原 conversation 和已接受 Turn;转换 owner 必须是显式、受 fence 保护的 handoff,不能由 enabled 或 updated_at 隐式完成。补充多会话、退出模式、关闭源会话与 lost-receipt 后换会话的回归,证明同一 intent 不能在另一 session 再次派发。

动机

这是 roadmap R3 的实际交付缺口:member result 已 accepted,协调者却必须自己轮询才能继续。让 Chat host 在既有开启模式、授权绑定和总额度内继续同一个 lead,是有用的有界产品增量。但 R3 的“回原会话、只触发一次、丢回执可恢复”必须成立,accepted、Turn 被保存和实际派发不能互相替代。上述两处故障正好落在该验收路径上,而不是额外要求完成所有宿主或云端生命周期。

改动思路

TS observation owner 在第一次 accepted transition 生成稳定 intent;Python 将它与 accepted 同写,in-Turn 读取则标记 observed_in_turn。host pump 复用 ChatLoopXMode 与原生 submit_turn,不增加新的 Goal 或额度;typed chat-mode planner 保留 mode、pause、绑定、原生状态与 allowance 的拒绝规则。这些 seams 和同写边界合理。缺口在 Python 的 recovery shortcut 和动态会话选择,绕开了已有 Turn dispatch 与原会话关系,而非 intent hash 本身。

具体改动

共 12 个文件,新增 847 行、删除 23 行:Chat wake host/pump、collaboration intent/receipt、两个现有 TS owner、耐久回归、双语文档和派生 census。完整 diff、旧 caller 和 related return/delegation contract 已阅读,没有把另一 PR 的修复当作本 head 的证据。

关键代码讲解

  • transitionDelegationObservation(delegation.ts:334)仅在首次 accepted transition 生成 requester/result 绑定 intent;accepted 重读不会产生第二条 intent,终态非 accepted 不触发。
  • Delegations._observe(collaboration_mcp.py:689)把 accepted 与 pending wake 原子记录;record_wake(316)在 operation 锁下结算,工具读取与 host 使用一致锁序。
  • ChatLoopXMode._wake_decision(chat_loopx_mode.py:473)先恢复 client Turn,再走 typed host admission。现在前半段把 queued 存在直接当 woken,是第一个阻塞点。
  • _select_owner(923)与 pump_delegation_wakes(950)发现 pending operation、挑选目标并隔离单条故障;重试未固定 session,形成第二个阻塞点。

本人重跑六个 Python suite 共 55 项、TS delegation/chat-mode 共 21 项、architecture 28 项均通过,typecheck 通过。相同 fixture 在不可变 base/head 通过真实 mode apply → controller → TS acceptance 与持久化 readback,普通 configure-off、start、completed-resume refusal 的完整规范化结果一致;disabled 不派发,普通 start 单次派发,拒绝全文及零副作用保持。仅规范化随机 ID、时间、临时 fixture 根目录及由该目录参与的 digest,并同时保留完整 config/request,未删除语义字段。另加独立负例 3 项均失败,覆盖 queued 恢复、原会话隔离和 lost-receipt 跨 session 重复派发。

对主干的风险

这是 enabled Goal Chat 的默认行为变化,不是新 actor 生命周期或跨 Agent 权限。普通单会话 off 路径的 paired parity 已验证,但源会话 disabled 后另一会话仍可消费它的 wake,故不能认定 scoped default-off 已隔离。实际执行仍要经过注册绑定、原生 admission 与预算 gate;execution guidance 是模型提示,不是完成/收养 obligation 的替代。状态和错误文字保持 domain-neutral,没有新增 substring classifier。后台 service 自带启动停止边界,未证明已安装 App 或 Lark 的运行体验。

语义与 CI 对齐

远端 CI 不查询、不轮询。风险 canary 执行 11 个 selected 检查和 5 个 direct 检查,唯一失败为现有 semantic twin budget 44 independently maintained py/ts twins; budget is 43。同一原始检查在 immutable base 和本 head 的失败身份/细节相同,PR 不改其因果路径,census 全量生成字节一致,错误行 mutation 仍被拒绝;该既有红项不决定评审。此次阻塞来自独立复现的本 PR 唤醒语义:queued 不是 dispatched,动态选中的同身份新会话不是原 conversation,也不能使 session-local client id 成为跨会话单次派发证明。

我的整体评价

这是有用的 R3 slice,但恢复和返回绑定未达到标题承诺。bounded future-facing pass:将 pump 留在现有 Chat mode owner 而非新增顶层模块是合理的;当前必须复用原生 dispatch 状态机和既有原会话返回关系,避免另建 replay 决策源。更大的 return/wake tick 合并可以继续 deferred,不作为本次修复前提。未运行 live model、packaged App 或 Lark,不能把 transport 替身扩展成这些证据。修复两个 blocker 后重新审查完整 exact head;control-plane 变更留给 maintainer 合并。

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
A persisted queued Turn is not a completed wake dispatch.  The wake decision
now carries the Turn already accepted under the intent's client id to the
typed planner: only a Turn that started is recorded as woken, a still-queued
Turn is replayed through the same native acceptance and dispatch owner, and a
Turn that ended before it started is refused instead of being reported as a
wake.  Pending intents therefore recover from an adapter failure between
acceptance and dispatch.

Signed-off-by: song <liusongstep@gmail.com>
planDelegationWake now refuses an intent whose conversation is not this
conversation, or whose coordinator is not the configured one, and it reads the
wake Turn's own facts: a started Turn is dispatch evidence, a queued one is
replayed under the same admission (including its own active Turn id), and a
Turn that ended before it started refuses the wake.  The planner returns
woken/admitted/refused/pending with the exact dispatch the host must perform,
so the host can no longer infer dispatch from a Turn's existence.

Signed-off-by: song <liusongstep@gmail.com>
The trusted Chat host records the session and Turn that started a delegated
operation, once, on first creation; a replay under the same operation id never
rebinds it, and the model supplies no routing.  The accepted-result wake intent
carries that conversation, so a wake returns to the conversation that owes it
and to no other.

Signed-off-by: song <liusongstep@gmail.com>
The originating conversation is part of the intent identity, so the same
accepted result cannot be woken into a different conversation of the same
requester.

Signed-off-by: song <liusongstep@gmail.com>
Real ChatRuntimeController.submit_turn, TS acceptance and durable store with
faults injected only at the adapter and worker transport: adapter failure after
acceptance is re-dispatched as the same Turn, a prepared capsule is repaired,
a lost receipt replays only the original Turn, a Turn cancelled before it
started is not woken, a queued Turn keeps the pause boundary, and neither
exiting, closing nor replacing the origin conversation moves the wake to
another session.

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

English verdict: both blockers addressed on the new exact head; please re-review.

Reviewed head: 3573ade2d7993e2c72b8b41741c3804ea52c6886 (upstream/main 3ec049e13 merged in, --signoff, no rebase). Both findings were reproduced by a failing test before the fix, on real ChatRuntimeController.submit_turn + TS acceptance + durable store, with faults injected only at the adapter/worker transport (no model, no provider).

[P1] queued Turn ≠ completed dispatch — fixed in 43fcdb834, 5e74c7beb

The decision no longer short-circuits on "a Turn exists". _wake_decision passes the Turn already accepted under the intent's client id to the typed planner (wake_turn: id, status, started_at, loopx_execution, operation, intent_id), and planDelegationWake decides:

  • started → woken (dispatch: "recorded"); dispatch is not repeated.
  • queued → admitted with dispatch: "replay", and the host re-enters submit_turn with the stored message and request, so chat_turn_acceptance.ts dispatchFor()/planSettledReplay/planPreparedRepair own the recovery. Same Turn, no second Turn. The wake's own queued Turn is excluded from the lead_turn_active check, so replay is not self-blocked.
  • ended before it started → refused: wake_turn_not_started (its client id cannot admit another Turn).

Evidence: test_adapter_failure_after_acceptance_is_redispatched_not_recorded (first tick: adapter init fails, Turn queued with started_at=None, 0 dispatches, intent stays pending; second tick: 2 adapter attempts, 1 dispatch, 1 wake Turn, receipt woken with created=False), test_failure_after_the_prepared_capsule_is_repaired_and_dispatched (prepared capsule → repaired → dispatched), test_lost_receipt_replays_only_the_original_turn[running|terminal], test_wake_turn_cancelled_before_it_started_is_not_woken, test_queued_wake_turn_keeps_the_pause_boundary (pause still refuses, zero dispatch).

[P1] conversation pinning — fixed in 03efe2150, be361feb(=be361c160), 5e74c7beb

The implicit selector is gone (_owners/_select_owner deleted). The trusted Chat host records conversation={session_id, turn_id} beside the operation on first creation only (Delegations.start(..., conversation=...), written by chat_loopx_mode.dispatch from its own session and Turn, never model-supplied), and delegationWakeIntent folds it into intent_id. The pump now wakes only that session: no conversation → refused: no_wake_owner; closed/exited/session missing → refused: no_wake_owner; another conversation or a reconfigured coordinator → refused: wake_identity_conflict. No implicit handoff, and the scheduler deadline recheck remains the fallback.

Evidence: test_wake_never_moves_to_another_conversation_after_exit, test_closed_origin_conversation_is_refused_without_a_substitute, test_intent_without_an_originating_conversation_wakes_nobody, test_pinned_conversation_rebound_to_another_requester_is_refused, test_lost_receipt_then_new_conversation_cannot_dispatch_twice (first tick accepts + dispatches in the origin session, receipt write lost; origin conversation exits; new same-identity conversation is enabled → retry lands in the origin session, created=False, exactly 1 dispatch, and the substitute session has no wake Turn), plus test_the_starting_conversation_is_pinned_beside_the_operation and test_the_in_turn_tool_pins_the_conversation_it_runs_in (a second start under the same operation id never rebinds the pin).

Verification on this head

  • Python: pytest tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py tests/test_chat_delegation_journey.py tests/test_collaboration_mcp.py tests/test_local_delegation.py tests/test_delegation_inventory.py tests/architecture/test_project_registry_io_census.py tests/architecture/test_top_level_module_budget.py → 81 passed. Wake suite alone: 22.
  • TS: node --test tests/control_plane_ts/{chat_mode,delegation,turn_acceptance}.test.ts → 22 passed, 0 failed; npm run typecheck:control-plane ok.
  • Mutation check (each fix reverted, its test fails, then restored): queued-Turn shortcut → 6 fail; not-started treated as dispatch evidence → 1 fail; conversation dropped from the planner identity → 1 fail; conversation dropped from intent_id → 1 fail; conversation pin not written by Delegations.start → 2 fail; pump ignoring the pin → 1 fail. All restored.
  • loopx canary premerge --from-git-diff --git-diff-base upstream/main: gate passed, tier standard, 5 direct + 11 selected checks, 0 failures, 0 manual holds.

Docs updated for both behavior changes (goal-chat-continuation.md, local-delegation.md, EN+zh).

Not run: a live model Turn started by a wake, the packaged App and Lark — unchanged from the previous head; the wake still goes through the existing submit_turn.

Control-plane change; leaving review and merge to the maintainer.

Review 修正已提交,新 head 3573ade2d:queued Turn 不再当成已派发(改为交回原生 acceptance/dispatch 恢复同一个 Turn,未启动就终止的 Turn 明确拒绝),唤醒按「启动该操作的会话」固定绑定,不再隐式选择其他同身份会话;两个阻塞点均有先失败后通过的回归与变异验证,请复审。

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST_CHANGES — 来源 Session 的固定与已有 queued 回合重放已改善,但新建/重放 dispatch 后仍会提前写终态 woken。native worker 的首次启动事实落盘失败时,Turn 没有启动,存储恢复及 controller 重建后 pump 却都不再重试。这是当前恢复语义的 P1,不是外部 CI 红灯。

Exact head: 5304@3573ade2d7993e2c72b8b41741c3804ea52c6886;比较基线:3ec049e138917a8cce4f84197ba196d26445b2b0。按 LoopX PR-review capability policy 12 审阅完整差异和当前修复反馈,而不是沿用旧头结论。

动机

Roadmap R3 所需的是接收方完成独立验收之后,原 lead 能继续下一次有界 Turn,不要求 owner 反复轮询。此前 delegation 进入 accepted 不会调度发起它的 Goal Chat lead;追加消息也不等于开始回合。此 PR 面向已激活、本地管理的 Chat lead,是合理的 R3 增量,不是整个多 Agent 协作或 Goal 完成。正常路径减少“结果已接受但无人继续”的停顿,用户继续使用既有 start/resume、pause/exit。但投递终态必须对应可恢复的启动事实,否则服务故障仍会把自动继续变成 owner 手动救援。

改动思路

首个 accepted 转换在现有 delegation journal 上记录 intent;可信 host 在开始 delegation 时固定来源 Session/Turn,intent 身份包含这个不可变来源。typed TypeScript owner 判断是否合法继续,Python 负责已有 Chat store、锁、native 接受和 dispatch 传输,不另建决策源。pump 是投递器而非权限授予者:结果被接受不代表可以唤醒任意会话,更不代表 adopted、Todo 完成或父 Goal 完成。已 started 的同一 intent 回合可补记事实;仅 queued 时应通过原 native 路径修复未完成 dispatch。当前 typed 分支遵守这一区分,但 Python 在 submit 返回后无条件终结 intent,破坏了同一条规则。

具体改动

完整差异为 13 文件、+1164/-21,集中于既有 Chat mode/runtime/server、delegation journal、两处 TS owner、语义登记、回归测试和公开操作文档。新增状态是不可重算的 accepted-result identity/origin 与投递回执,不是第二套 canonical settlement。Session 与 journal 锁顺序一致,accepted replay 不重新发 intent,已有回合观察到结果时也不会再强行唤醒。

关键代码讲解

  • transitionDelegationObservation / delegationWakeIntent 只在首次进入 accepted、canonical completion 和 artifact 当前时产生 intent;requester、artifact identity 与可信来源共同确定身份,accepted→accepted 不重新调度。
  • planDelegationWake 验证原 Session、Goal、注册 lead、绑定、启用状态和剩余额度。其他活跃 Turn、pause 或预算耗尽保留 pending;关闭/解绑/原 Goal 完成拒绝。只有同 intent 的真正 started 回合记为 woken;queued 回合走 replay,不凭存在就宣称启动。
  • ChatLoopXMode.wake / _wake_decision 将 typed 决定送入原 submit_turn/native acceptance 和 worker 路径,稳定的 intent client id 把补偿锁在同一回合,而不是抢另一条 Session。不过 L569 无条件 _woken(...),没有验证异步 worker 的启动事实已持久化。
  • pump_delegation_wakes 从 canonical execution records 发现未终结 intent,校验 storage address,隔离坏记录;每 tick 限制的是实际改变的回执,不让前面 unchanged pending 记录遮住后面的可处理结果。它由现有 Chat server 托管并随之停止。

对主干的风险

[P1] 不要在启动事实尚未持久化时终结 wake intent。 位置 chat_loopx_mode.py:569。用真实 submit_turn、_start_accepted_turn_worker 和 _run_turn,只在隔离 Chat store 的第一次 queued→starting 写入注入一次 OSError。观察到 Turn 仍 queued、started_at=null,journal 却已 woken。恢复原 store writer 后再次 pump,随后用同一持久化 store/registry 重建真实 controller 再 pump,两次均返回空结果,原 Turn 仍未启动:record_wake 和 _pending_identity 都只接受 pending,终态回执使意图从恢复扫描永久消失。

这与只在“已有 Turn”分支区分 queued/started 不同:submit_turn 返回只证明接受及异步 thread dispatch,并不保证启动事实已落盘。最小修复应由现有 native dispatch/receipt owner 保留可恢复状态,直到同 intent 的 started 事实可读回;并让首次启动写入失败正确释放 worker single-flight 状态,以便恢复后继续同一个 Turn,而不是新建 Turn、换 Session 或提高额度。回归必须覆盖首次启动写入失败/进程中断、存储恢复、controller 重建、同 client id 重试、一次实际启动和之后终态去重。

来源改绑问题已验证修复:host 固定 conversation,模型参数不能写它,找不到原 owner 不会自动改绑。此前已 queued、adapter dispatch 尚未完成的重放和 worker single-flight 测试也通过;它们不能替代上述更晚的失败窗口。普通 Chat mode-off 的配置、start、完成 Goal 的 resume 拒绝与基线保持一致,包括完整返回、错误、配置、持久化和 dispatch 观察,而非只比较 decision code。

独立执行了 81 项 Python 测试、22 项 TS 测试、control-plane typecheck 和标准 premerge(5 direct + 11 selected),均通过;基线对应现有 Python 57 项、TS 19 项也通过。额外真实 Chat controller/store 对照的 3 个普通入口观测完全相同(标准化摘要 f02408b7130f9e5d71b7f80b7b09fe204571432373f7c7e8d5245924b);只去除随机 id、时刻及合成根路径派生摘要,保留字段、错误、持久化、消息和副作用。另两个独立检查覆盖真实 worker 的 queued replay 单次启动,以及 25 条 unchanged paused pending 前置后仍能处理后续记录,均通过。新增上述启动事实失败的独立检查则失败,准确暴露 premature terminal receipt。前两个检查的初版 fixture 有时序断言和 reason 名称错误,修正后重跑通过;故障检查初次有 import 路径设置错误,正确设置后才取得产品反例。未修改 PR 源码、放宽检查或把评审 setup 错误当成产品失败。

实际 mode/controller/store 和 journal 运行在隔离合成 Goal 中;付费模型/远端 delegated process 使用受控 transport substitute,未认证全部署 soak、远程 peer 生命周期或完整 R3 接收采纳旅程。pump 仍扫描 execution records;本次只证明发现完整性及变化上限,不宣称大规模长期扫描性能已达标。这是明确的残余验证边界,不应靠扩大超时或无关框架冒充解决。

语义与 CI 对齐

延伸现有 delegation observation、Chat mode/native acceptance 词汇,登记了真实生产 call site,没有新建跨 Agent authority。wake 是 host-only operation,不进入普通 web owner operation;只有已有显式启用的受管理 Goal Chat lead 可以被继续。公开文档明确“只有真正启动过的 Turn 才算已派发”,typed owner 也用 started 区分事实;当前 Python 终态写回与之不一致,不能只修改文案把 queued 说成 started。这里的 guidance 只是既有结果提示;入场、限额、身份是机器规则。修复后重跑 uv run --extra test python -m pytest -q tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py,扩展真实 worker/重建恢复用例,再跑 TS owner 测试和对应 premerge。未查询、轮询或等待远程 CI。

我的整体评价

长期推进判断为“正常路径改善、故障恢复仍有缺口”,用户体验为“启用后的回传改善、关闭路径保留,但故障后仍可能要手动救援”;不把一条 woken 回执当作全系统持续运行证明。新增可靠性状态有真实 producer,不能仅靠当前 UI 状态重算;保留现有 TS 决策 owner 和 Python 传输边界是合适的,但终态必须受启动事实约束。未来向检查考虑了 session 选择、worker single-flight 和 journal 扫描:原 owner 已足够,不需要第二个唤醒框架;扫描优化若有实际规模证据再定位既有 owner,不为本 PR 造任务。完整范围与原问题相称、关闭路径兼容也已验证,但本 PR 自身的恢复义务尚未满足,因此要求修改;修复后重新核验新精确头。本次未合并。

English verdict: REQUEST_CHANGES — The current head fixes requester-session identity and earlier queued replay, but terminalizes wake after asynchronous submit without durable start evidence. A first worker-start write failure leaves the Turn queued and the wake woken; storage recovery and a fresh controller cannot rediscover it. Keep the same intent recoverable until matching start evidence, with native-worker cleanup and restart coverage.

…ct lands

`submit_turn` returning proves admission and an asynchronous dispatch, not
that the Turn's start fact is durable.  The wake path wrote a terminal
`woken` receipt right after it returned, so a first worker-start write
failure left the Turn queued with the intent already settled: storage
recovery and a freshly built controller both stopped rediscovering it.

Two changes, in their owning modules:

- `_wake_decision` now records `pending` / `wake_dispatch_pending` after
  admission and lets the next tick read the durable `started_at` back, so
  the receipt stays a statement about a start that actually persisted.
  The typed owner distinguishes `started_at` from a status guess, keeps a
  Turn that is still activating pending, and refuses only a Turn that
  ended without starting, reusing `isTerminalTurnStatus` rather than
  restating the terminal set.
- `_run_turn` released its single-flight guard only when `worker.start()`
  itself raised.  A failure writing `queued -> starting` escaped the
  worker body and left the key in `turn_done_events` forever, so every
  later dispatch of that Turn was a silent no-op.  The start write and
  the Turn body now run under one `finally` that hands the guard back.

Coverage pins the failure window itself: a one-shot `OSError` on the
first `queued -> starting` write, then the same Turn replaying through
native dispatch to one start, a rebuilt controller over the persisted
store rediscovering the intent, and terminal de-duplication after that.
Characterization of the pass path is unchanged.

Signed-off-by: song <liusongstep@gmail.com>
…-lead

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Re-review request — exact head b76fb6993efeddeb85b9b7dededc3a589fb7db0d

Thanks for the concrete failure window. I reproduced it on the reviewed head first, fixed the two causes, then re-ran the same fault path.

[P1] A terminal woken receipt before the start fact landed — fixed

chat_loopx_mode.py:569 wrote _woken(...) unconditionally after submit_turn returned. Your reproduction showed the consequence: the first queued -> starting write raises OSError, the Turn stays queued with started_at=null, and the intent is already terminal — so record_wake / _pending_identity only accept pending and both storage recovery and a freshly built controller stop rediscovering it.

Fix. _wake_decision now records settled("pending", "wake_dispatch_pending", turn_id=...) after admission. The receipt becomes a statement about a start that actually persisted: the next tick reads the durable started_at back through the same typed decision and records woken then. The typed owner distinguishes started_at from a status guess, keeps a Turn that is still activating pending, and refuses only a Turn that ended without starting — reusing isTerminalTurnStatus from chat_turn_acceptance.ts rather than restating the terminal set beside it.

The second cause: the worker single-flight guard was never released

Your review noted this should be handled by the existing native dispatch owner. While reproducing, the retry stayed refused even after the store was healthy, which the receipt path alone cannot explain:

  • _start_accepted_turn_worker registers turn_done_events[key] before the start fact is durable, and cleared it only when worker.start() itself raised.
  • _run_turn's queued -> starting write sits outside any try/finally in the worker body. An OSError there escaped and left the key in turn_done_events permanently, so every later dispatch of that Turn was a silent no-op.

Fix (chat_runtime.py): the start write and the Turn body now run under one try/finally that hands the guard back, so recovery continues the same Turn through the existing path rather than creating a new one, moving Session, or raising a limit.

Evidence

tests/test_chat_delegation_wake.py now covers the failure window itself, over the real submit_turn / _start_accepted_turn_worker / _run_turn with faults injected only at the worker transport and the store write — no model, no provider:

  • a one-shot OSError on the first queued -> starting write: the Turn stays queued, the receipt stays pending with no woken_at, and the single-flight key is released;
  • the next tick replays that same Turn through native dispatch to exactly one start, then a later tick reads the start fact back and records woken with created=false;
  • a rebuilt ChatRuntimeController over the same persisted store rediscovers the intent and starts the same Turn once;
  • terminal de-duplication: the following tick is a no-op.

Verification on the merged head: tests/test_chat_delegation_wake.py + tests/test_chat_loopx_mode.py 42 passed; chat_mode.test.ts 5 passed; control-plane typecheck clean. Latest origin/main merged with sign-off (--no-ff, no force-push).

The remaining CI red on this head is inherited from main, not from this diff: test_prompt_upgrade_hook.py::test_live_decision_adds_only_existing_required_read_channel[loopx_turn_run_once] and Execution chip is not a compact hairline row: 28px tall both reproduce on current main and are being fixed on their own PRs.

English verdict request: the intent now stays recoverable until matching start evidence, with the worker guard released, and both the same-process and rebuilt-controller paths are covered. Please re-review b76fb6993efeddeb85b9b7dededc3a589fb7db0d.

@cocolord cocolord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

评审提交:b76fb6993efeddeb85b9b7dededc3a589fb7db0d。这是 policy-12 whole-PR、exact-head 复审。我重新核对了 GitHub 的完整 14 文件差异(+1424/-36)、提交与 checks,复用了上一轮在未变更边界上的证据并检查了失效条件,重点验证 3573ade2d 之后对 queued replay、首次 start 写失败和 worker single-flight 的修复。此前的 conversation pin 与 queued Turn 误判已实质修好;当前仍有一个更靠后的可靠性 blocker、一个公开协议漂移,以及一个由本 PR 引入的 required maintainability gate 失败。

动机

这个 PR 解决的用户问题明确且值得做:成员结果已经通过独立验收并进入 accepted 后,空闲的 Goal Chat LoopX 协调员此前不会自动继续,只能碰巧在自己的 Turn 中轮询,或等待 owner 再次操作。新行为限定在已显式开启、未暂停、绑定仍有效、原生 Goal 可恢复且额度充足的 managed Codex Goal Chat;它不触碰非 Chat/Lark/attached-host,不解除暂停,不新开 Goal,也不把 accepted 当作 adopted、Todo 完成或父 Goal 完成。正常路径能减少“结果已有、主线不动”的长期停顿,收益是真实且可观察的。

改动思路

首次进入 accepted 时,typed delegation owner 生成包含 requester、operation/request、artifact digests 和可信来源 conversation 的 intent;Python 在同一 requester-scoped operation record 中原子保存 wake: pending。Chat server 托管一个轻量 pump,只扫描 pending intent,固定回到启动 delegation 的原 Session,再由 planDelegationWake 复用现有 resume 的 Goal、身份、binding、pause、active Turn 和 allowance 规则。真正执行仍走 ChatRuntimeController.submit_turn 与 native acceptance;intent 派生的稳定 client id 让 prepared/queued 丢回执能够恢复同一个 Turn,而不是创建第二个 Turn。lead 在自己的 Turn 内 read/wait/adopt 时则写 observed_in_turn,避免多余唤醒。

这个 owner 切分总体正确:TS 决策、Python effect、现有 Chat store/native Turn、server lifecycle 都被复用,没有再造 scheduler 或跨 Agent authority。当前问题不在总体方向,而在 started_at 的含义跨层被放大:Python 把它写成 worker 开始处理的事实,TS 却把它解释成 provider 已实际接收 Turn 的终态证据。

具体改动

完整差异覆盖两份公开文档;chat_loopx_mode.py 的 guidance、host-only wake、pending 扫描与服务;collaboration_mcp.py 的 execution path、receipt、conversation pin 与 in-Turn suppression;chat_server.py 的启动/关闭;chat_runtime.py 的 wake display 和 worker guard;三处 typed contract;registry I/O manifest;以及四组 Python/TS 回归。测试占新增行的大部分,但生产边界仍增加了新的 durable intent/receipt 和周期性扫描语义。

关键代码讲解

  • transitionDelegationObservation / delegationWakeIntent 只在首次合法进入 accepted 时产生 intent;accepted→accepted 不重复调度,conversation 进入 identity,避免同 Goal/同 Agent 的另一会话接管。
  • Delegations._observe / record_wake / wake_observed_in_turn 把结果与 wake 分成不同事实,并在同一 record lock 下从 pending 走向 woken/refused/observed;因此任何 terminal receipt 都必须真实,因为后续不会再进入恢复扫描。
  • pump_delegation_wakes / ChatLoopXMode._wake_decision 校验 storage address 与 pinned Session,隔离单条坏记录,调用 typed owner,并把 queued/prepared 同 client id 交回既有 native acceptance 修复。当前 head 已把 submit_turn 返回后的回执从无条件 woken 改为 wake_dispatch_pending。
  • planDelegationWake 集中处理 identity、mode、Goal、binding、active Turn、native 状态和 allowance;也复用了 shared terminal Turn status。但 L94 把任意 started_at 直接当作 dispatch: recorded。
  • ChatRuntimeController._run_turn / _run_started_turn 新的外层 finally 正确修复了第一次 start 写失败时 single-flight key 泄漏;不过 started_at 在进入 _run_started_turn 之前就写入,后面仍要做 context、mode、driver 和 adapter/provider dispatch。真正的上游启动事件会在 _TurnEventBuffer 中把 Turn 改为 running 并记录 upstream_turn_id。

对主干的风险

[P1] started_at 不是 provider dispatch 证据,当前会产生不可恢复的假 woken。 chat_mode.ts:94 先看 started_at 就返回 woken,甚至早于 terminal status 分支;而 chat_runtime.py:1406 在 prepare_turn_context、loopx_mode.prepare、CodexGoalDriver.run / adapter.start_turn 之前就写入该字段。独立 exact-head 探针走 shipped store/controller/planner/receipt 两条路径:一条在写入后模拟进程退出并重建 controller;另一条让 prepare_turn_context 在 provider 调用前抛错并由真实错误路径把 Turn 标为 failed。两条都没有 turn.started 事件、没有 provider dispatch,但下一 tick 都把同一 intent 写成 terminal woken。之后 _pending_identity / record_wake 不再处理它,自动继续永久丢失。

最小修复应复用现有真实上游启动事实(例如 durable turn.started 对应的 running / upstream_turn_id),或新增一个明确但同 owner 的 dispatch receipt;对 pre-dispatch starting、failed、server_restarted 分别定义可恢复或可操作的终态。回归必须覆盖“start 写成功、provider 调用前失败/进程退出、重建 store/controller、不得 woken、同 intent 最终只实际启动一次、之后去重”,不能让 mock 直接把被验证的 postcondition 当作 dispatch。

[P2] 公开 reason vocabulary 与实现不一致。 当前 EN/ZH 文档仍列 wake_turn_not_started,而实现/测试已经改为 wake_turn_ended_unstarted;新增的 pending wake_dispatch_pending 也未列入文档。这里是 operation readback 的机器状态,不只是解释性文案。请确定稳定名称、同步两种语言,并用契约测试防止再次漂移。

Required gate 也由本 PR 引入失败。 精确 base 7e60e6999 的 maintainability ratchet 通过;精确 head 因 module_metric_budget:loopx/chat_runtime.py 未登记增长而失败。当前 loopx canary premerge --from-git-diff 共选 11 项,10 项通过,唯一失败就是该 catalog canary。请优先把新增职责放到最近的有界 owner;若维护者明确接受 hot-module 增长,再按仓库机制更新 reviewed ceiling 并说明收益。不能把 required red 仅归为 main 噪声。

其余验证:focused Python 42 passed;focused TS 22 passed;control-plane typecheck 和 changed-Python Ruff 通过;public boundary 与其他 9 个 selected canary/smoke 通过。远端当前为 19 success、11 failure、7 skipped;日志中的 28px UI、prompt-upgrade/generated-twin 等失败至少有主干既有成分,但 aggregate gate 仍红,我没有用这些噪声替代上述本 PR 反例。未运行付费 live-model、打包 App/Lark、跨平台 hard-kill 和大规模 journal 扫描性能,这些保留为残余验证边界。

语义与 CI 对齐

这个 PR 是对既有 delegation observation、Chat mode 与 native Turn 词汇的合理扩展;identity/permission 规则是 typed 且 domain-neutral,没有 substring denylist,也没有把强制 gate 写成 guidance。真正不对齐的是:公开文档声称“只有真正启动过的 Turn 才算已派发”,而 current typed rule 使用的是更早的 worker timestamp;同时文档枚举与实现 reason 不同。请修复事实边界和文档,再重跑 pre-dispatch fault/restart、focused Python/TS、typecheck、Ruff、reason contract、maintainability ratchet 与 premerge。

我的整体评价

结论是 REQUEST_CHANGES。功能动机明确,正常路径对长期推进和用户体验明显正向;conversation pin、queued replay、首次 start-write 失败和 single-flight cleanup 都已得到实质修复,现有 owner 选择也合理。但当前 exact head 仍会把“worker 准备开始”记录成“lead 已被唤醒”,这正好破坏本 PR 最核心的恢复承诺;公共 reason 也已漂移,且 required maintainability gate 是 head 新增失败。请在现有 planner/runtime owner 内做有界修复,不需要继续扩展框架;新 head 到来后我会重跑这两个 pre-provider 反例、reason 对齐和 exact-base/head gate。本次只发布 review,不修改作者代码,也不授予 merge authority。

English verdict: REQUEST_CHANGES on exact head b76fb6993efeddeb85b9b7dededc3a589fb7db0d. The feature has clear positive value and the prior conversation-pin, queued-replay, first start-write, and single-flight issues are materially fixed. However, started_at is persisted before context/provider dispatch, yet the planner treats it as terminal wake evidence; two real-store/controller fault probes produced woken with no turn.started event. Align success with durable provider-start evidence, synchronize the public reason vocabulary, and clear the PR-attributable chat_runtime.py maintainability ratchet before re-review.

if (turn.loopx_execution !== true || turn.operation !== "wake" || turn.intent_id !== intent.intent_id) {
return outcome("refused", "wake_identity_conflict");
}
if (typeof turn.started_at === "string" && turn.started_at) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] started_at cannot certify this terminal woken outcome. _run_turn persists it before _run_started_turn, so a process loss or prepare_turn_context failure can leave no turn.started event and no provider dispatch. On this exact head, both cases are nevertheless finalized as woken on the next pump, after which pending-intent recovery no longer sees them. Please base terminal success on the existing durable provider-start evidence (or an equivalent explicit dispatch receipt), define the pre-dispatch failed/restarted outcome, and cover controller rebuild plus exactly-one eventual start.

`pending` with `lead_turn_active`, `lead_paused` or `allowance_exhausted`, or
`refused` with `goal_stopped`, `lead_unbound`, `binding_revoked`,
`native_goal_complete`, `native_goal_absent`, `wake_identity_conflict`,
`wake_turn_not_started` or `no_wake_owner`. Its Turn is accepted once under a

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] The public reason list is stale on this head: the typed owner/tests now emit wake_turn_ended_unstarted, while this EN/ZH document still advertises wake_turn_not_started; the new pending reason wake_dispatch_pending is also omitted. Since these values are exposed by delegation readback, please synchronize both language sections and add a small contract check for the emitted reason set.

The planner read `started_at` as dispatch evidence, but the worker stamps that
field before it builds the turn context, prepares LoopX mode and hands the
message to the adapter. A Turn that failed in that window — or whose process
exited before the provider call — was therefore recorded as `woken` although the
provider never accepted it. A terminal receipt is never rescanned, so the intent
was lost and the lead was never continued.

Dispatch is now proven by `upstream_turn_id`, which the runtime writes when the
provider reports `turn.started` and checkpoints immediately. A Turn carried past
its start write but not dispatched stays `wake_dispatch_pending`; one that ended
without dispatch is `refused` / `wake_turn_ended_unstarted`. The Python owner
passes the provider fact through and no longer records a terminal `woken` of its
own after `submit_turn` returns, so the typed decision is the only writer of
that state.

The public reference now lists the reasons the planner actually produces —
`wake_dispatch_pending` added, `wake_turn_not_started` replaced by
`wake_turn_ended_unstarted` — in both languages, and two contract tests pin the
vocabulary to the planner so a rename cannot drift from the docs again.

`chat_runtime.py` had grown past its reviewed `any_count` ceiling. The growth was
the attachment payload annotation on the extracted start body, so the payload is
named once as `AttachmentPayload` and the four call sites share it; the module is
now below the ceiling rather than above it, and the extracted body no longer
takes a `done_event` its caller already owns.

Coverage: a worker start with no provider dispatch stays pending across a tick
and a rebuilt controller, and only the provider's own fact wakes it once; a Turn
that failed before dispatch is refused and mints no second Turn.

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Re-review request — exact head 4b5b3d44b

Both probes reproduced on the reviewed head before the fix. You were right that the failure boundary was later than the one I had fixed.

[P1] started_at is not provider dispatch evidence — fixed

_run_turn stamps started_at when it moves the Turn queued -> starting, before prepare_turn_context, loopx_mode.prepare and the adapter call. Reading it in the planner recorded woken for a Turn the provider never accepted, and since a terminal receipt is never rescanned, the intent was lost.

Dispatch is now proven by upstream_turn_id, which _TurnEventBuffer writes when the provider reports turn.started (and checkpoints immediately). The planner uses it and nothing earlier:

  • dispatched → woken / dispatch: recorded, no second Turn;
  • past its start write, still activating, or never reached the provider → pending / wake_dispatch_pending;
  • ended without dispatch → refused / wake_turn_ended_unstarted.

_wake_decision no longer writes a terminal woken of its own after submit_turn returns, so the typed owner is the only writer of that state. Both paths you described are covered: the process-exit case stays pending across a tick and a rebuilt controller over the persisted store, and the pre-provider failure ends in an actionable refusal; neither mints a second Turn, and only the provider's own fact wakes once.

Mutation-checked in both directions: restoring the started_at read fails the TS contract, and passing started_at as upstream_turn_id fails three Python cases.

[P2] Reason vocabulary drifted from the implementation — fixed

The reference now lists what the planner actually produces — wake_dispatch_pending added, wake_turn_not_started replaced by wake_turn_ended_unstarted — in both languages. Two contract tests pin it: every documented reason must appear as an outcome(...) in the planner, and every planner reason must appear in both documented lists, so a rename cannot drift again without failing.

The maintainability ratchet is cleared, not raised

The growth was the attachment payload annotation on the extracted start body. Rather than raising the reviewed ceiling, the payload is named once as AttachmentPayload and the four call sites share it; any_count is now 32, below the 35 ceiling, and the extracted body no longer takes a done_event its caller already owns. tests/canary/test_maintainability_ratchet.py passes.

Verification on this head

tests/test_chat_delegation_wake.py + test_chat_loopx_mode.py + test_chat_store_input_validation.py + test_chat_codex_home.py 51 passed; chat_mode.test.ts 5 passed; ruff clean; ratchet 14 passed; loopx canary premerge --from-git-diff 8 risk-profile smokes and the public-boundary scan pass with 0 failures.

English verdict request: success is now aligned with durable provider-start evidence, a pre-dispatch Turn stays recoverable instead of recording a false woken, the public vocabulary matches the planner, and the required gate is green without a ceiling change. Please re-review 4b5b3d44b.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact head: 5304@4b5b3d44ba7b402ff7bd84245a19ccce82d1e20e; immutable base: 7e60e69999d9dd00c6d19a93221982a2f1f736b1。本次按 LoopX PR-review capability policy 12 重新执行 whole-PR review,不继承旧头结论,也未查询或等待 GitHub CI。

动机

成员结果已经独立验收,而发起工作的 Goal Chat lead 仍然空闲,这是值得解决的协作停顿。PR 希望在原会话显式启用 LoopX、未暂停、绑定及 Goal 仍有效、额度允许时继续一次 Turn。这个结果是 roadmap R3 的有界增量,不是整个跨宿主协作验收或父 Goal 完成。正常唤醒能改善持续推进与用户轮询负担,但必须保留未启用 Chat 唤醒的 CLI/MCP 委派路径;当前完整差异还没有满足这个隔离边界。

改动思路

实现复用原生 delegation acceptance、Chat Turn admission/client id、会话锁、operation 锁与 provider 启动事实。接受结果产生意图,Chat service 的 pump 找到发起工作的固定会话,再由既有模式 owner 判断 pending、refused 或 woken;quota、pause、Goal 生命周期与独立验收没有第二套 owner。固定会话与 accepted result 身份是有意义的持久事实,provider 的 upstream_turn_id 也比 worker 的 started_at 更可靠。

不过,明确的原会话/启用状态必须在产生 capability 状态之前限定范围,不能先给所有普通委派创建 wake,再用 refused 补救。更小的修复是在现有 typed transition 与 host producer 之间明确“有唤醒资格的 Chat-origin result”,保留普通 acceptance 的旧输入及读回契约,不新增一个平行 scheduler 或万能开关。

具体改动

完整 14 文件包括 Chat 模式与 runtime、服务启动、共享 collaboration adapter、三个 TypeScript owner、registry IO manifest、两份参考文档及 Python/TS 回归。相对上一评审头 b76fb6993efeddeb85b9b7dededc3a589fb7db0d,本头把 woken 判断移到 provider 启动事实,维护 queued/replay 身份,并修正启动失败时的 single-flight 收尾及公开原因词汇;这些修复已通过本次回归,不继续沿用旧 blocker。

关键代码讲解

  1. transitionDelegationObservation(delegation.ts:334)仍验证 canonical completion、acceptance 和当前 artifacts,但第 341 行现在对每次首次 accepted 无条件调用 delegationWakeIntent。后者要求 requester/result 身份,却允许 conversation 为 null。这是关闭态漂移的源头,不是后台 pump 单独造成的。
  2. Delegations._observe / _wake_requester(collaboration_mcp.py:699–718)把 typed intent 写入 operation,并从 host 的 conversation 关系取路由。正常 CLI/MCP acceptance 也经过这个共享调用;随后 public wait/read 可带上 wake,因此不能把它视为仅 Chat 的内部诊断字段。
  3. ChatLoopXMode.wake(chat_loopx_mode.py:453)在固定会话与 operation 锁内复用原生接受、原 client id 和原 queued Turn。未派发、adapter 失败、暂停与额度不足保持可恢复 pending;只有 provider 启动事实才写 woken。重建 controller 后仍重放原回合,没有借重试新建第二个回合。
  4. pump_delegation_wakes(chat_loopx_mode.py:953)由 Chat server 启动,每次只处理有限 changed intents。没有原会话的 intent 写为 no_wake_owner/refused;退出或关闭原会话不会改选同 requester 的另一会话。后者防止错投,但前者仍是在未激活范围写 capability 状态。

对主干的风险

[P2,需修复的契约问题] 普通非 Chat 委派被纳入 Chat wake 的状态与读回协议

位置:delegation.ts:341,连到 _observe 和全 runtime pump。

独立反例使用相同普通委派任务、同一原生验收规则和 fixture model transport,不创建 Chat 会话,也不启用 Chat LoopX 模式。在 File 与 SQLite 两种真实 authority 上,base 都是 accepted、canonical Todo done、一次 Host invocation、无 wake,journal 在空 Chat pump 前后不变。head 的相同任务仍 accepted 且只执行一次,但 public wait 已新增 wake;journal 有 pending wake,调用生产 pump 后变为 refused/no_wake_owner,journal 被改写。base 2/2 通过,head 2/2 未满足独立的关闭态不变条件。

这不是已证明的越权模型启动:反例没有第二个 Host invocation。问题是未激活路径的共享输入/持久投影/public readback 被改了,关闭態隔离不能由“最终没启动模型”替代。typed accepted transition 的旧合法参数也新增了 wake requester 依赖。最低修复是只为明确授权的 Chat-origin operation 产生并消费 wake intent,普通 accepted transition 保持兼容;对真正启用的原会话仍保留一次投递、暂停、额度、失效绑定与重启恢复。

本次本地验证:106 项 Python 回归通过,22 项相关 TS 测试通过;npm run typecheck:control-plane、改动路径 ruff 均通过。独立 source premerge 的 5 项 direct、2 项 catalog、8 项 risk 与 1 项 public/private boundary 全通过;这不覆盖上面的独立反例,也不是 merge 权限。所有执行使用 exact-head checkout 环境;未调用付费模型,未更改运行中的 Goal。

语义与 CI 对齐

当前 obligation 是 opt-in capability 在关闭范围不得改变共享持久投影、输入要求和用户读回。PR 扩展既有 delegation/wake vocabulary,而不是创建更广的 actor 生命周期;typed 状态、domain-neutral 错误、guidance 与 enforced admission 的区分总体清楚。违规点是这个共享 owner 的无条件 intent producer。复审应同时跑普通 File/SQLite start → native validation → wait → empty-Chat pump 的 base/head 对照,以及 tests/test_chat_delegation_wake.py、tests/test_local_delegation.py 和相关 TS 测试。此旧头的 semantic inventory 不支持新的 --changed-from advisory 参数;已用其支持的完整报告及 premerge full-tree drift 检查,未把工具版本限制当成 PR defect。CI 按 capability policy 不参与本次判定。

我的整体评价

English verdict: REQUEST_CHANGES — exact head 4b5b3d44ba7b402ff7bd84245a19ccce82d1e20e; ordinary non-Chat delegation gains wake state and public readback despite no activation, reproduced on real File/SQLite authority with a passing immutable-base control.

前几轮的 dispatch/restart 修复有实质价值,机制规模与问题基本相称;不会因为代码量、作者身份或未知 CI 拒绝。持续推进的正向路径已改善,但未激活用户的 readback/持久状态出现明确回归,因此 user_experience 判为 regression,long_horizon 的启用路径有改善、跨关闭范围保持未完成。未来面向重构检查选择“在现有 typed acceptance owner 收窄 producer”这一小修复;不要求广泛 TS 重写或另立 framework。既有 operation/client-id 的持久重放义务应保留,不删真实兼容 seam。修复隔离并加入上述对照后可复审;App/Lark、跨宿主与付费 provider 未验证,也不声称本地测试覆盖它们。当前不建议合并。

…gin result

An ordinary CLI/MCP delegation was gaining wake state it had no use for. The
typed accepted transition produced an intent for every first acceptance and let
it carry a null conversation, so a plain `wait`/`read` grew a `wake` field, the
record gained a persisted pending intent, and the runtime pump rewrote it to
refused/no_wake_owner — on both File and SQLite authority, for a caller who never
created a conversation or enabled this capability. Producing an intent and then
refusing it is not opt-in isolation: the shared persistent projection and the
public readback changed for everyone.

The requester's conversation is the capability's own precondition, so it is now
also its producer's: an accepted result leaves a wake intent only when the
operation was started from a conversation, which only the trusted Chat host ever
supplies. A conversationless acceptance keeps exactly the transition it had —
no intent, no wake state, and no change to what a plain read returns — while the
enabled path keeps its single delivery, pause, allowance, binding and
restart-recovery behaviour.

`test_an_ordinary_delegation_gains_no_wake_state` pins the isolation on both
authorities: acceptance succeeds, no `wake` and no `conversation` appear on the
record. Mutation-checked by restoring the unconditional producer, which fails it.
The TS contract states both directions — a conversationless acceptance has no
intent, a conversation-origin one is pinned to that conversation — and the
forged-digest assertions now travel on the path that actually validates them.

Signed-off-by: song <liusongstep@gmail.com>
…-lead

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Re-review request — exact head e60074dd0

Reproduced your control on both authorities before fixing: an ordinary CLI/MCP delegation that never created a conversation or enabled the capability still gained a persisted pending wake, a wake field on public wait, and a journal rewrite to refused/no_wake_owner once the pump ran. The problem was the producer, exactly as you located it.

[P2] Opt-in isolation: the producer now requires its own precondition

The requester's conversation is what makes a wake meaningful, and only the trusted Chat host ever supplies it — so it is now the producer's precondition too, not something the pump repairs afterwards:

if (to === "accepted" && from !== "accepted" && wakesItsConversation(params)) {
  return {status: to, wake_intent: delegationWakeIntent(params)};
}
return {status: to};

wakesItsConversation reads the requester's own conversation. A conversationless acceptance therefore keeps exactly the transition it had: no intent, no wake state on the record, and no change to what a plain wait/read returns. The enabled path is untouched — single delivery, pause, allowance, revoked binding and restart recovery all remain as reviewed.

Evidence

test_an_ordinary_delegation_gains_no_wake_state runs the ordinary start → native validation → wait path on File and SQLite and asserts acceptance still succeeds while no wake and no conversation appear on the record. Mutation-checked by restoring the unconditional producer, which fails both.

On the TS side the contract now states both directions: a conversationless acceptance yields {status: "accepted"} with no wake_intent, and a conversation-origin acceptance is pinned to that exact conversation, so the intent identity still distinguishes two conversations of the same requester. The forged-digest/goal-ref assertions moved onto the path that actually validates them, since a conversationless acceptance no longer reaches that code.

I did not add a second eligibility source: row.conversation is written once by Delegations.start from the host-supplied argument and never replaced, so the Python side already carries the same fact, and the intent's identity hash still includes it.

Verification on this head

tests/test_chat_delegation_wake.py + test_chat_loopx_mode.py + test_local_delegation.py 66 passed; delegation.test.ts 17 passed; latest origin/main merged with sign-off. The App/Lark, cross-host and paid-provider journeys remain unverified as you noted.

English verdict request: an operation outside a conversation no longer gains wake state, persistent projection or readback, while the enabled conversation path keeps its reviewed behaviour and recovery. Please re-review e60074dd0.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] 当前 Chat server 登记锚点仍漂移,必需检查需同步

Reviewed exact head: e60074dd0b13b25182238de6e409a3a6517f8d25; immutable merge base: 67930ab6af78491f10ca3de4ff74ef7a39954a51.

动机

当前切片解决local Chat lead在委派结果accepted之后仍需手动poll/续跑的缺口,归属roadmap R3结果回流。目标是原请求会话一次可恢复的native continuation,不是registered peer授权、Goal完成,也不是付费provider+App/Lark两轮长链路的全部验收。

改动思路

可信in-turn host在delegation start钉住session/turn;TS canonical accepted transition原子生成accepted-bound intent,Python journal负责持久化;background pump按现有mode/native binding/budget计划,复用native submit_turn与client-id,不另造scheduler或decision owner。wake不会unpause、启动native Goal、扩大预算或转移到同Goal的另一会话。

具体改动

完整14文件涵盖四个Python host模块、三个TS owner、两篇public reference、manifest与native/TS验证。此前普通File/SQLite delegation因为缺少requester而拒绝acceptance的问题已修复:transitionDelegationObservation 只对有可信origin的工作生成wake;同一base/head真实start→fixturehost→native验证→wait→canonical done/readback,两种store均保持accepted、一次host调用、相同8个普通read字段、wake缺席,unactivated journal没有变更。

关键代码:delegation.ts:335 transitionDelegationObservation 仍由canonical accepted驱动;chat_mode.ts:80 planDelegationWake 根据 upstream_turn_id 而非started_at认定provider已启动,queued/starting pending、terminal unstarted refused;chat_loopx_mode.py:453 wake 按session→operation锁调用原native Turn;chat_runtime.py:1387 _run_turn 的外层finally现在覆盖第一笔queued→starting写入,瞬时失败也释放single-flight;chat_loopx_mode.py:953 pump_delegation_wakes 扫operation journal,按发生变化的receipt计limit,service随serve_chat启动/停止。

我还自己运行真实native worker:第一次start写入注入一次OSError,普通/重建controller两种情况都释放并重放原Turn;随后fixture transport在provider前结束,真实状态为failed,wake正确refused为wake_turn_ended_unstarted,重复pump不新建第二Turn、不冒称woken。此负向证据不等于付费provider送达;正向/暂停/关闭/重绑/重复dispatch由完整native和TS fixtures验证,边界明确。

当前需修P2是 project_registry_io_manifest_v1.json 的三个current-source锚点。ChatRequestHandler._goal_channel_extension_ready 登记966但源码968,serve_chat 登记1488但源码1490,新 serve_chat._wake_goal_context 登记1576但源码1578。test_checked_in_project_registry_io_manifest_is_current 与 test_semantic_vocabulary_registry_matches_the_code 在当前头真实失败,semantic canary同样报告这三个站点;相同两套architecture文件在不可变base为123pass。因此不是“非自身PR引起红CI”,也不重复已经解决的旧wake阻塞。

最小修复是用已有 uv run --extra test python scripts/generate_project_registry_io_manifest.py 同步登记,查看classification没有改变,再跑这两个architecture suite与风险型premerge。不要降低必需检查、另建manifest、回退有效wake逻辑或扩大到无关迁移。

对主干的风险

本轮selected native为203pass/2fail(上述登记/semantic),TS44pass,typecheck、6个changedPython Ruff、配置内19-source Mypy、advisory、diff check通过。5项premerge direct通过;11项selected中semantic1项失败,其余catalog/risk8/public boundary通过。初始aggregate同时缺少本轮CQA;我已记录自己的非通过CQA,未将缺失reviewer步骤伪装为author运行时缺陷。第一版review命令错误(不存在test路径、额外probe import)停在collection,已在同一未改source上纠正并保留失败回执,不算本PR问题。

默认关闭态采用真实File/SQLite base/head对照,不以“enabled path没跑”推导隔离。可信origin、session/Goal/agent/ref及budget是实际effect gate;安装、accepted artifact、同Goal其他conversation本身不授权wake。新intent是原请求的不可推导事实,woken是native dispatch的派生receipt,不能代替接收采纳或canonical Goal验收。旧无origin rows仍可读/可接受,不需要平行legacy决策源。全部新增vocabulary限定delegation wake,通用错误保持goal-neutral;execution guidance是建议,native admission和receipt条件是机器执行约束。

没有frontend导航/settings/editor第一屏变化,server创建/运行service就是此host切片入口;未跑installed frontend/Lark/真实收费provider,不声称完成父R3。当前阶段是可独立测试、回滚的local host增量,下一步沿用原roadmap验收,不新增仪式性跟进任务。

wait_for_ci=false:未查询、轮询或等待remote CI。REQUEST_CHANGES仅依据同一检查在immutable base通过而exact head失败的本地归因,其他红CI不影响评审结论。

我的整体评价

REQUEST_CHANGES,范围仅这一条P2生成登记同步。普通关闭态、原会话绑定、provider证据及初始写失败cleanup的旧问题已得到当前头验证,不要求重做。未来维护性pass确认复用native admission/typed owner和既有生成器,未发现需要本批额外框架或宽迁移。请同步三个锚点后对新exact head重跑必需检查与CQA/premerge;runtime/control-plane仍由maintainer决定合并,评审不是自合并授权。

English verdict: REQUEST_CHANGES - the wake and ordinary-delegation fixes validate, but three changed Chat server registry I/O anchors remain stale. The identical immutable-base checks pass; regenerate the existing manifest and rerun required native checks on the new head.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants